Papers with open-source toolkit

35 papers
PhoNLP: A joint multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and dependency parsing (2021.naacl-demos)

Copied to clipboard

Challenge: PhoNLP is a multi-task learning model for joint Vietnamese part-of-speech (POS) tagging, named entity recognition (NER) and dependency parsing.
Approach: They propose a multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and dependency parsing that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently.
Outcome: The proposed model outperforms a single-task learning approach that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently.
Reliable, Reproducible, and Really Fast Leaderboards with Evalica (2025.coling-demos)

Copied to clipboard

Challenge: Using open-source evaluation tools, we create reliable and reproducible model leaderboards with human and machine feedback.
Approach: They propose an open-source evaluation toolkit that facilitates the creation of reliable and reproducible model leaderboards.
Outcome: The evaluation tool facilitates the creation of reliable and reproducible model leaderboards.
NeurST: Neural Speech Translation Toolkit (2021.acl-demo)

Copied to clipboard

Challenge: a toolkit for speech translation is available for free and provides step-by-step recipes for feature extraction, data preprocessing, distributed training, and evaluation.
Approach: They propose to use NeurST to facilitate speech translation research for NLP researchers . they show experimental results for different benchmark datasets which can be regarded as reliable baselines .
Outcome: The proposed framework provides reliable benchmarks for speech translation research.
MarkLLM: An Open-Source Toolkit for LLM Watermarking (2024.emnlp-demo)

Copied to clipboard

Challenge: Large Language Models (LLMs) embed imperceptible yet algorithmically detectable signals in outputs to identify LLM-generated text.
Approach: They propose to develop an open-source toolkit for LLM watermarking that embeds imperceptible yet algorithmically detectable signals in model outputs to identify LLM-generated text.
Outcome: MarkLLM provides a unified framework for implementing LLM watermarking algorithms, while providing user-friendly interfaces to ensure ease of access.
OpenSLU: A Unified, Modularized, and Extensible Toolkit for Spoken Language Understanding (2023.acl-demo)

Copied to clipboard

Challenge: Spoken Language Understanding (SLU) is a task-oriented dialogue system . open-source toolkit provides a unified, modularized, and extensible toolkit for SLU .
Approach: They introduce an open-source toolkit to provide a unified toolkit for spoken language understanding.
Outcome: The proposed toolkit unifies 10 models for both single-intent and multi-intention scenarios.
LeafNATS: An Open-Source Toolkit and Live Demo System for Neural Abstractive Text Summarization (N19-4)

Copied to clipboard

Challenge: Neural abstractive text summarization (NATS) has gained a lot of attention in the past few years from both industry and academia.
Approach: They propose an open-source toolkit for training and evaluation of different sequence-to-sequence based models for the NATS task and for deploying the pre-trained models to real-world applications.
Outcome: The proposed model can be used to generate high-quality summaries that are verbally innovative and can easily incorporate external knowledge.
fastHan: A BERT-based Multi-Task Toolkit for Chinese NLP (2021.acl-demo)

Copied to clipboard

Challenge: Recently, the need for Chinese natural language processing (NLP) has a dramatic increase for many downstream applications.
Approach: They propose to use Chinese word segmentation (CWS), Part-of-Speech (POS) tagging, named entity recognition (NER), and dependency parsing to train a multi-task model based on a pruned BERT.
Outcome: The proposed model performs better than popular segmentation tools on a non-training corpus.
LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems (2024.acl-demos)

Copied to clipboard

Challenge: Existing tools for augmented question-answering do not support researchers and developers to customize the training, testing, and deployment process.
Approach: They propose an open-source toolkit that features a wide selection of model training algorithms, evaluation methods, and deployment tools curated from the latest research.
Outcome: The proposed framework trains and deploys 7B-models with the same performance as OpenAI’s text-ada-002 and GPT-4-turbo.
RiskLab: A Controlled Toolkit for Probing Emergent Risks in LLM-Based Multi-Agent Systems (2026.acl-demo)

Copied to clipboard

Challenge: Recent advances in large language model (LLM) agents have accelerated deployment of multi-agent systems for complex tasks.
Approach: They propose an open-source toolkit for instantiating, probing, and measuring emergent risks in LLM-based multi-agent systems under controlled conditions.
Outcome: The proposed toolkit is based on a structured topology–environment–protocol–agent–task quintuple enabling reproducible studies of how communication structure, coordination mechanisms, and incentives shape system-level risks.
TokLens: A Multilingual Lens on Tokenizer Quality for LLMs (2026.acl-srw)

Copied to clipboard

Challenge: TokLens is an open-source toolkit for evaluating tokenizer quality across languages . authors evaluated 24 tokenizers from major LLM families across 15 typologically diverse languages - a gap that is stark in Japanese .
Approach: They evaluate 24 tokenizers from major LLM families across 15 typologically diverse languages and correlate these metrics with downstream performance.
Outcome: The proposed tokenizers produce 56x more tokens per word in Japanese than in English . the newer tokenizer Qwen2.5 and Gemma-2 reduce this gap to under 4x .
AutoAlign: Get Your LLM Aligned with Minimal Annotations (2025.acl-demo)

Copied to clipboard

Challenge: Automated Alignment (ALM) is a set of algorithms designed to align Large Language Models (LLMs) with human intentions and values while minimizing manual intervention.
Approach: They propose an open-source toolkit that integrates mainstream automated algorithms through a consistent interface and an accessible workflow supporting one-click execution for prompt synthesis and automatic alignment signal construction.
Outcome: The proposed framework enables easy reproduction of existing results through extensive benchmarks and facilitates the development of novel approaches via modular components.
ConvLab-2: An Open-Source Toolkit for Building, Evaluating, and Diagnosing Dialogue Systems (2020.acl-demos)

Copied to clipboard

Challenge: ConvLab-2 inherits Convlab's framework but integrates more powerful dialogue models and supports more datasets.
Approach: They present ConvLab-2, an open-source toolkit that enables researchers to build task-oriented dialogue systems with state-of-the-art models and perform an end-to-end evaluation.
Outcome: The new tool inherits ConvLab's framework and extends it by integrating many recently proposed state-of-the-art dialogue models.
NeuSpell: A Neural Spelling Correction Toolkit (2020.emnlp-demos)

Copied to clipboard

Challenge: a new spelling correction toolkit is available for free.
Approach: They propose an open-source toolkit for spelling correction in English . they train neural models using spelling errors in context and using richer contextual representations.
Outcome: The proposed spell-checker improves accuracy on synthetic examples and richer representations of the context.
NeuroX Library for Neuron Analysis of Deep NLP Models (2023.acl-demo)

Copied to clipboard

Challenge: NeuroX is an open-source toolkit to conduct neuron analysis of natural language processing models.
Approach: They propose a Python toolkit to conduct neuron analysis of natural language processing models.
Outcome: a new open-source toolkit enables neuron analysis of natural language processing models . the framework provides a framework for data processing and evaluation, making it easier for researchers and practitioners to perform neuron analyses.
CRSLab: An Open-Source Toolkit for Building Conversational Recommender System (2021.acl-demo)

Copied to clipboard

Challenge: Existing studies on conversational recommender systems lack a unified and standardized implementation or comparison.
Approach: They propose to use a unified framework and highly-decoupled modules to develop CRSs.
Outcome: The proposed framework collects 6 commonly used human-annotated CRS datasets and implements 19 models that include advanced techniques such as graph neural networks and pre-training models.
BMInf: An Efficient Toolkit for Big Model Inference and Tuning (2022.acl-demo)

Copied to clipboard

Challenge: Recent years, pre-trained language models (PLMs) have achieved promising results on various NLP tasks.
Approach: They propose an open-source toolkit for big model inference and tuning which can support big model tuning at extremely low computation cost.
Outcome: The proposed toolkit can support big model inference and tuning at extremely low computation cost.
RobustQA: A Framework for Adversarial Text Generation Analysis on Question Answering Systems (2023.emnlp-demo)

Copied to clipboard

Challenge: Question answering (QA) systems have reached human-level accuracy, but they are not robust enough and vulnerable to adversarial examples.
Approach: They modified the attack algorithms widely used in text classification to fit them for QA systems.
Outcome: The proposed framework is the first open-source toolkit for investigating textual adversarial attacks in QA systems.
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation (2026.acl-demo)

Copied to clipboard

Challenge: Existing document translation pipelines face a tension between linguistic processing and layout preservation.
Approach: They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content.
Outcome: The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision.
Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation (P19-3)

Copied to clipboard

Challenge: Texar is an open-source text generation toolkit that supports a broad set of text generation tasks.
Approach: They introduce Texar, an open-source text generation toolkit that supports text generation tasks.
Outcome: Texar supports machine translation, summarization, dialog, content manipulation, and more.
OpenT2T: An Open-Source Toolkit for Table-to-Text Generation (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing methods for table-to-text generation are limited and benchmarked on a limited number of datasets.
Approach: They propose to use open-source tools to reproduce existing large language models for performance comparison and expedite the development of new models.
Outcome: The proposed toolkit compares existing large language models on 9 table-to-text generation datasets and maintains a leaderboard to provide insights for future work.
Open-Theatre: An Open-Source Toolkit for LLM-based Interactive Drama (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing tools for creating, modifying, and experimenting with interactive dramas are limited.
Approach: They propose an open-source toolkit for creating configurable LLM-based interactive drama.
Outcome: The proposed toolkit enhances narrative coherence and realistic behavior in interactions with agents.
PyOpenDial: A Python-based Domain-Independent Toolkit for Developing Spoken Dialogue Systems with Probabilistic Rules (D19-3)

Copied to clipboard

Challenge: a recent development of spoken dialogue systems has enabled deep learning to achieve state-of-the-art performance.
Approach: They propose a Python-based domain-independent, open-source toolkit for spoken dialogue systems.
Outcome: The proposed toolkit extends OpenDial's Java-based architecture and provides new functions for neural dialogue state tracking and action planning.
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation (2024.emnlp-demo)

Copied to clipboard

Challenge: Large language models (LLMs) lack domain-specific knowledge and can cause hallucinations.
Approach: They propose a RAG diagnosis tool that visualizes the attentiveness of the generated tokens in retrieved documents.
Outcome: RAGViz provides token and document-level attention visualization and generation comparison upon context document addition and removal.
LEGOEval: An Open-Source Toolkit for Dialogue System Evaluation via Crowdsourcing (2021.acl-demo)

Copied to clipboard

Challenge: Currently, researchers use automatic metrics and human evaluation to evaluate dialogue systems.
Approach: They propose to use a Python API to easily evaluate dialogue systems using Amazon Mechanical Turk.
Outcome: The open-source toolkit provides a fast, consistent method for reproducing human evaluation results.
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails (2023.emnlp-demo)

Copied to clipboard

Challenge: NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Approach: They propose to add programmable guardrails to LLMs that are user-defined, independent of the underlying LLM, and interpretable.
Outcome: The proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails.
KMatrix-2: A Comprehensive Heterogeneous Knowledge Collaborative Enhancement Toolkit for Large Language Model (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing studies on K-LLMs systems focus on declarative knowledge and procedural knowledge (rules) .
Approach: They propose to build a toolkit that supports comprehensive heterogeneous knowledge collaborative enhancement for Large Language Models (LLMs).
Outcome: The proposed toolkit provides unified knowledge integration and joint knowledge retrieval methods to achieve more comprehensive heterogeneous knowledge collaborative enhancement.
OpenICL: An Open-Source Framework for In-context Learning (2023.acl-demo)

Copied to clipboard

Challenge: In-context Learning (ICL) is a new paradigm for large language model evaluation.
Approach: They propose an open-source toolkit for ICL and LLM evaluation.
Outcome: The proposed framework is highly flexible and flexible and can be easily combined with other tools to suit users' needs.
WIKIR: A Python Toolkit for Building a Large-scale Wikipedia-based English Information Retrieval Dataset (2020.lrec-1)

Copied to clipboard

Challenge: ad-hoc information retrieval methods usually require large amounts of annotated data to be effective.
Approach: They propose an open-source toolkit to automatically build large-scale English information retrieval datasets based on Wikipedia.
Outcome: The proposed toolkit builds large-scale English information retrieval datasets based on Wikipedia with 59,252 queries and 2,617,003 pairs.
CSPB: Conversational Speech Processing Benchmark for Self-supervised Speech Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on clean, single-speaker, single channel audio, failing to reflect the complexities of natural human interaction.
Approach: They propose a benchmark to assess the robustness of self-supervised speech models in conversational settings.
Outcome: The proposed benchmark assesses the robustness of self-supervised speech models in conversational scenarios.
Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question Answering (C18-1)

Copied to clipboard

Challenge: Sentence pair modeling is a fundamental technique underlying many NLP tasks.
Approach: They analyze several neural network designs for sentence pair modeling and compare their performance extensively across eight datasets.
Outcome: The proposed models perform well across eight datasets including paraphrase identification, semantic textual similarity, natural language inference, and question answering tasks.
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pretrained Models (2026.findings-eacl)

Copied to clipboard

Challenge: Extending existing vocabulary is a widely used step in adapting pre-trained language models to new domains or languages.
Approach: They propose to extend a pre-trained tokenizer by continuing the BPE merge learning process on new data.
Outcome: The proposed method improves tokenization efficiency and improves model utilization.
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) encode and store contextual information, but internal mechanisms are opaque.
Approach: They propose a toolkit that assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions and measures intrinsic dimensionality of representations.
Outcome: The proposed framework assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions, and measures the intrinsic dimensionality of representations.
On the Evaluation of Speech Foundation Models for Spoken Language Understanding (2024.findings-acl)

Copied to clipboard

Challenge: Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech.
Approach: They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs .
Outcome: The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks .
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on how SAEs derive most fine-grained latent features for safety remain unexplored.
Approach: They propose a framework for interpreting SAE features in safety-critical domains . they train a suite of SAEs with human-readable explanations and systematic evaluations based on pornography, politics, violence, and terror .
Outcome: The proposed framework reduces interpretation cost by 55% and improves safety-critical features.
skLEP: A Slovak General Language Understanding Benchmark (2025.findings-acl)

Copied to clipboard

Challenge: skLEP is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding models.
Approach: They introduce a benchmark specifically designed for evaluating Slovak natural language understanding models.
Outcome: The proposed benchmark covers nine tasks that span token-level, sentence-pair, document-level tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations